Papers by Xuan Long Do

7 papers
Aligning Large Language Models with Human Opinions through Persona Selection and Value–Belief–Norm Reasoning (2025.coling-main)

Copied to clipboard

Challenge: Current methods for reasoning and predicting human opinions employ role-playing with personae but face two major issues: LLMs are sensitive to even a single irrelevant persona, skewing predictions by up to 30%; and LLM fail to reason strategically over personas.
Approach: They propose a four-step solution modeling which and how to reason with personae, inspired by the Value–Belief–Norm theory.
Outcome: The proposed model improves existing methods by up to 4% by fine-tuning them with COO's data.
UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for chart-based data analysis neglect explicit modeling of chart structures.
Approach: They propose a pretrained model for chart comprehension and reasoning that encodes relevant text, data, and visual elements of charts and uses a chart-grounded text decoder for text generation.
Outcome: The proposed model outperforms existing methods that lack explicit modeling of chart structures and lacks explicit modeling.
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning (2022.findings-acl)

Copied to clipboard

Challenge: Existing datasets that focus on complex reasoning questions do not address such questions as they are template-based and answers come from a fixed-vocabulary.
Approach: They propose a large-scale benchmark that uses visual and logical reasoning to answer questions using a transformer-based model.
Outcome: The proposed models achieve state-of-the-art on the previous datasets and on the current one, but also show that they have several challenges in answering complex reasoning questions.
Modeling What-to-ask and How-to-ask for Answer-unaware Conversational Question Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to generate conversational question are naive and do not account for the answer span.
Approach: They propose a framework for generating a conversational question from a context.
Outcome: The proposed framework achieves state-of-the-art in two different settings compared to existing models . it uses a sentence as the rationale and extracts the answer span from it .
OpenCQA: Open-ended Question Answering with Charts (2022.emnlp-main)

Copied to clipboard

Challenge: OpenCQA is a task to answer open-ended questions about charts with descriptive texts.
Approach: They propose a task to answer open-ended questions about charts with descriptive texts.
Outcome: The proposed task is to answer an open-ended question about a chart with descriptive texts.
CoHS-CQG: Context and History Selection for Conversational Question Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing studies focus on single-turn question generation, but few studies have studied the challenges of multiturn QG.
Approach: They propose a two-stage conversational question generation framework that shortens the context and history of the input and calculates relevance scores.
Outcome: The proposed framework achieves state-of-the-art on CoQA in answer-aware and answer-unaware settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations